Back

Scientific Data

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match Scientific Data's content profile, based on 209 papers previously published here. The average preprint has a 0.15% match score for this journal, so anything above that is already an above-average fit.

1
Hybrid transcriptome assembly and annotation of Japanese macaque prefrontal cortex

Chatzipli, A.; Voshall, A.; Viswanadham, V.; Weiss, A. R.; Liguore, W. A.; McBride, J. L.; Sherman, L. S.; Lee, E. A.; Yu, T. W.

2026-08-19 neuroscience 10.64898/2026.08.10.743719 medRxiv
Top 0.1%
56.2%
Show abstract

Japanese macaque (Macaca fuscata) is used in biomedical and neurobiology research, yet transcriptomic resources for the brain are limited. We present a hybrid RNA sequencing dataset and a prefrontal cortex transcriptome assembly from two healthy 6-year-old animals. Short-read Illumina ({approx}70 million paired-end reads per sample) and long-read Oxford Nanopore direct RNA sequencing ({approx}2.5 million reads per sample) were combined. Reads were quality controlled, aligned to the macFus_1.0 reference genome, and assembled with StringTie2. Transcripts were annotated using Trinotate and eggNOG-mapper, and open reading frames were predicted with TransDecoder. The released data package includes raw reads (NCBI SRA BioProject PRJNA1295993), transcript sequences and structural annotation files, predicted coding sequences and proteins, functional annotation tables, and transcript abundance estimates (TPM). Technical validation includes read-level QC and protein-level comparisons to expressed gene sets from human, rhesus macaque and chimpanzee prefrontal cortex. These resources enable reuse for transcript-level expression studies, isoform characterization and comparative primate neurogenomics.

2
A versioned, analysis-ready archive of United States State Cancer Profiles county- and state-level estimates

Davis, S.

2026-08-27 health informatics 10.64898/2026.08.24.26361254 medRxiv
Top 0.1%
46.5%
Show abstract

State Cancer Profiles (statecancerprofiles.cancer.gov), maintained by the National Cancer Institute with the Centers for Disease Control and Prevention, is a widely used source of county- and state-level cancer statistics in the United States, used for cancer-center catchment-area surveillance and for geographic studies of cancer burden, screening, and access to care. The site offers no API, no bulk download, and no archive of prior estimates: its sole machine-readable export returns one statistical stratum per HTTP request, and when the underlying data are updated the previous estimates are overwritten and become unrecoverable. This resource provides the complete national county- and state-level extract of all four State Cancer Profiles data topics (incidence, mortality, screening and risk factors, and demographics) as typed, analysis-ready files with the stratifying dimensions as columns, published under pinned, citable version DOIs on Zenodo (concept DOI 10.5281/zenodo.11098814). One version DOI is minted per distinct upstream data vintage, the set of values the site served between successive replacements. Three vintages have been captured to date; at each observed vintage boundary roughly 97% of estimate values changed, so which vintage an analysis draws on affects its results. From the 2026-08-24 release forward, cells that the upstream site suppresses are retained as typed nulls with an explicit suppression-reason column. Capture has been automated on an approximately monthly cadence since February 2025, and each future upstream revision will be preserved as a new vintage.

3
Bridging Biomedical Atlas Ecosystem: Cross-Atlas Alignment And Scalable Tissue Specimen Registration

Jain, Y.; Desai, B.; Qaurooni, D.; Bhavsar, A.; Kienle, P.; Pouch, A. M.; ONeill, K.; Apte, S.; Herr, B. W.; Fisher, S. A.; Börner, K.

2026-08-22 bioinformatics 10.64898/2026.08.13.744704 medRxiv
Top 0.1%
31.7%
Show abstract

Over the last five years, over 13,000 tissue datasets with 200+ million cells from 20 consortia have been spatially registered into the Human Reference Atlas (HRA) common coordinate framework (CCF). The shared 3D spatial and semantic reference system enables exploration of datasets in the context of all other data across organs, assay types, and spatial scales. However, manual registration of individual samples remains resource intensive, posing feasibility challenges exacerbated by the proliferation of samples, assays, and atlasing efforts. This paper presents two approaches to scale up HRA construction: (1) projecting data across biomedical reference atlas systems and (2) using millitomes to bulk register tissue blocks into a reference organ. Both methods use the AutoMated Alignment and Projection (AMAP) pipeline to align 3D mesh models using point cloud registration. We demonstrate the evolving HRA-aligned atlas ecosystem for 6 models from the SPARC Program (heart), Gut Cell Atlas (large intestine), 500-subject consensus kidneys, and the Julich Brain Atlas. Additionally, we used AMAP to project 7 millitome models across 5 organs onto the HRA ecosystem, integrating 300+ tissue extraction sites. AMAP enables scalable tissue registration of data across atlas ecosystems enabling the construction of detailed reference maps of the human body.

4
A behaviourally normed database of 1,377 natural sounds for auditory cognition and neuroscience

Plegat, M.; Araujo Vitoria, M.; Marinato, G.; Tita, B.; van der Lans, C.; Pijfers, M.; Esposito, M.; Bertovic, M.-S.; Formisano, E.; Giordano, B. L.

2026-08-28 neuroscience 10.64898/2026.08.25.746933 medRxiv
Top 0.1%
31.7%
Show abstract

Natural-sound research requires stimulus sets that combine acoustic standardization with detailed behavioural characterization. We present 1,377 two-second sounds representing 240 expert-defined source--action classes. We call this database "MaMa Sounds", as it resulted from the collaborative effort of two academic teams in Maastricht and Marseille. The sounds were manually curated, segmented, sampled at 16 kHz, and labelled with a noun identifying the source and a verb identifying the action. We release deidentified trial-level identification and familiarity data together with multiple per-sound norms (e.g., identification accuracy, confidence and agreement; familiarity), along with overall norms derived with principal component analysis. Noun, verb, and joint noun--verb norms are provided as direct means and medians with the number of contributing observations. This battery preserves process-specific information, while two principal-component scores provide compact overall behavioural-identifiability measures derived from response ease, semantic correspondence, agreement, and familiarity. The repository also contains deterministic response-cleaning code, participant and reference Word2Vec representations, and code reproducing the public sound-level tables. The resource supports stimulus selection, matching, and continuous modelling in auditory cognition and neuroscience.

5
Chromosome-scale genome assembly and annotation of the Vietnamese indica rice cultivar Khang Dan 18

Nguyen, T. Q.; Do, K. H. D.; Vu, T. M.; Hoang, N. V.

2026-08-21 plant biology 10.64898/2026.08.15.742683 medRxiv
Top 0.1%
31.0%
Show abstract

Khang Dan 18 (KD18) is an Oryza sativa L. subsp. indica rice cultivar widely cultivated in northern Vietnam and used as an experimental and breeding background in Vietnamese rice research. Although KD18 has previously been represented in low-depth population resequencing datasets, a contiguous and annotated cultivar-specific genome has not been available. Here, we report a chromosome-scale genome assembly of KD18 generated using Oxford Nanopore long-read and Illumina short-read sequencing. The 395.3-Mb assembly comprises 12 chromosome-scale pseudomolecules containing approximately 95% of the assembled sequence and 99.6% of the predicted protein-coding genes. The assembly showed 97.2% BUSCO completeness, an average Merqury quality value of 46 and a long terminal repeat assembly index of 13.21. A total of 56,546 protein-coding genes representing 71,237 transcripts were predicted, with 99% BUSCO and 98.68% OMArk completeness. These statistics are similar to those of other high-quality genome assemblies that were recently published for different Asian rice cultivars, therefore providing a cultivar-specific genomic resource for research involving KD18 and KD18-derived materials.

6
Haplotype-resolved chromosome-level genome assembly of four European white oak species

Magris, G.; Avanzi, C.; Bagnoli, F.; Duvaux, L.; Belmonte, E.; Vendramin, G. G.; Piotti, A.; Pinosio, S.

2026-08-24 genomics 10.64898/2026.08.20.745905 medRxiv
Top 0.1%
27.1%
Show abstract

European white oaks (Quercus section Quercus) are ecologically and economically important forest trees characterized by extensive shared genetic variation and a history of interspecific gene flow. Genomic resources remain uneven across species, limiting comparative analyses and pangenome development. Here, we present haplotype-resolved chromosome-scale genome assemblies and genome annotations for four European white oak species: Quercus robur, Q. petraea, Q. pubescens, and Q. frainetto. The assemblies were generated from PacBio HiFi sequencing data and include both phased haplotypes for each species. Genome sizes range from 779 to 817 Mb and all assemblies are organized into 12 chromosome-scale pseudomolecules with high completeness and contiguity. We additionally provide species-specific repeat annotations, structurally and functionally annotated protein-coding gene sets, and complete organellar genomes. The dataset includes the first reference genomes for Q. pubescens and Q. frainetto, together with newly generated assemblies for Q. robur and Q. petraea produced using a consistent sequencing and analysis workflow. These resources provide a standardized framework for comparative genomics, pangenome construction, genome evolution studies, and investigations of adaptation and introgression across European white oaks.

7
WIO-ReefFish: A High-Resolution Dataset for Taxon-Aware Coral Reef Fish Detection in the Western Indian Ocean

Gerard, J.; Branger, L.; Huyghe, F.; Kochzius, M.; Otwoma, L.; Bergacker, S.; op't Roodt, L.; Rumisha, c.; Di Bella, L.

2026-08-20 ecology 10.64898/2026.08.19.745797 medRxiv
Top 0.1%
22.2%
Show abstract

Coral reef fish assemblages are widely used as indicators of ecosystem condition, yet manual annotation of underwater video remains a major bottleneck for scalable biodiversity monitoring. Despite rapid progress in automated detection, ecologically realistic and publicly available datasets remain scarce, particularly for the Western Indian Ocean. Here, we present WIO-ReefFish, a reef fish detection dataset derived from diver-operated line-intercept transects and designed for ecological monitoring under natural survey conditions. WIO-ReefFish comprises 1,000 ultra-high-definition images (3840 $\times$ 2160 pixels) and 6,768 exhaustive bounding-box annotations spanning 24 taxonomic categories, thereby preserving full-frame assemblage structure in complex reef scenes. We also establish a standardized benchmark across nine object detection models under two complementary protocols: class-aware detection and class-agnostic fish localization. Detection performance was consistently higher under the class-agnostic protocol. The best-performing model (RT-DETR) improved from 0.48 mAP50 in the class-aware setting to 0.70 mAP50 when taxonomic constraints were removed, indicating that taxonomic discrimination remains substantially more challenging than fish localisation in reef imagery. Spatially independent evaluation revealed a pronounced generalisation gap, particularly for taxonomic detection, whereas class-agnostic fish localisation remained substantially more robust across transects and countries. Together, these results establish WIO-ReefFish as a realistic benchmark for automated reef fish detection and provide a foundation for more robust computer-vision tools in coral reef biodiversity monitoring. The WIO-ReefFish dataset and associated benchmarking resources are publicly available.

8
A chromosome-level genome of the franciscana dolphin, Pontoporia blainvillei

Canesin, L. E. D.; Aleixo, A.; Vidal, A.; Martins, A. B.; Farro, A. P. C.; Kolesnikovas, C. K. M.; Cordeiro, D. d. M.; Neuhaus, E. B.; Luna, F. O.; Araujo, F. A. A.; Nunes, G.; Cunha, H. A.; Mendes, I. S.; Mattos, J. S.; Albuquerque, L.; Magalhaes, L.; Oliveira, R. R. M.; Bonatto, S. L.; Barreto, S. B.; Kantek, D. L. Z.; Vilaca, S. T.

2026-08-06 genomics 10.64898/2026.07.31.742074 medRxiv
Top 0.1%
18.9%
Show abstract

The franciscana dolphin (Pontoporia blainvillei) is a small coastal cetacean endemic to the southwestern Atlantic Ocean and one of the most threatened marine mammals worldwide. It faces severe threats from bycatch, habitat degradation, and pollution. Classified as "Vulnerable" by the IUCN and "Critically Endangered" in Brazil, the species restricted range, strong fidelity to shallow waters, and low reproductive rate increase its extinction risk. Here, we present the first chromosome-level genome assembly for the franciscana dolphin, generated using PacBio HiFi long-read sequencing and Hi-C chromatin conformation capture. The final assembly totaled 3.13 Gb across 22 chromosomes (1500 scaffolds), consistent with the estimated karyotype of 2n = 44, with scaffold N50 of 111.18 Mb, high BUSCO completeness (99.42%), and a consensus quality value of 65.76. This high-quality genomic resource fills an important phylogenetic gap within Cetacea, enabling comparative and conservation studies. It provides an essential foundation for population genomics research to assess genetic diversity, structure, and connectivity, thereby supporting evidence-based conservation strategies for this endangered species.

9
From Lab Notes to Linked Data: MeSyTo for Ontology-Driven Metadata in Toxicological Omics

Pozhidaeva, M.; Schreiber, S.; Schubert, K.; Busch, W.; Hackermüller, J.; Canzler, S.

2026-08-21 bioinformatics 10.64898/2026.08.14.736375 medRxiv
Top 0.1%
18.2%
Show abstract

Toxicological omics studies require comprehensive metadata to support reproducibility, interoperability, and regulatory reuse. However, metadata requirements differ across public repositories, reporting frameworks, and laboratory workflows, resulting in inconsistent annotation and limited data integration. To address this challenge, we developed MeSyTo (Metadata for Systems Toxicology), an ontology-driven framework for harmonizing metadata across toxicological omics. Metadata concepts from public repositories, the OECD Omics Reporting Framework (OORF), community standards, and institutional workflows were semantically aligned and implemented as the MeSyTo Metadata Model (MMM). The MMM serves as the basis for the automatic generation of SHACL validation shapes and framework-specific metadata profiles, while curated value sets are represented as SKOS controlled vocabularies to support metadata collection and validation. The current implementation comprises 105 ontology classes and 527 data properties and supports transcriptomics, proteomics, and metabolomics. A prototype web application demonstrates ontology-driven metadata collection with integrated semantic validation and ontology-based term resolution. The ontology, validation shapes, controlled vocabularies, generation scripts, and software are publicly available as open-source resources. MeSyTo provides a reusable semantic foundation for harmonized, machine-actionable metadata and facilitates repository submission, regulatory reporting, and interoperable data exchange across toxicological omics studies.

10
Next-generation insect digitization: combining phenomics and genomics by subsequent synchrotron X-ray imaging and DNA sequencing

Lupascu-Vasilita, C.; Riedel, A.; Mera-Rodriguez, D.; Cecilia, A.; Farago, T.; Hamann, E.; Hein, J.; Herz, A.; Martin, J.; Odar, J.; Pfeiffer, P.; Sarkar, C.; Spiecker, R.; Tavakoli, C.; Zuber, M.; Rabeling, C.; Baumbach, T.; Krogmann, L.; van de Kamp, T.

2026-08-24 genetics 10.64898/2026.08.20.745929 medRxiv
Top 0.1%
15.5%
Show abstract

Recent technological advances allow for the large-scale acquisition of genetic and morphological data: high-throughput sequencing has transformed the field of genomics while synchrotron X-ray microtomography enables rapid, noninvasive 3D imaging. However, integrating these approaches for the same specimens is challenging because X-rays can fragment DNA, and DNA extraction damages internal morphology, particularly relevant for small bodied organisms, such as insects. We systematically tested multiple extraction protocols and irradiation conditions across three model insect species. We irradiated more than 1,000 specimens under varying conditions and tested DNA quality through DNA barcoding and UCE sequencing. Our results demonstrate that high-quality DNA and high-resolution tomograms can be obtained from the same individuals, provided that the parameters are carefully optimized and rapid SR-CT scanning precedes DNA extraction. In this respect, our findings establish practical guidelines for combining genomics and phenomics, paving the way for comprehensive integrative digitization of biodiversity.

11
Gene model for the ortholog of Ilp3 in Drosophila pseudoobscura

Lieser, B. C.; Laskowski, L. F.; Huber, R.; Kolker, K. O.; Arsham, A. M.; Rele, C. P.; Toering Peters, S.

2026-08-23 genomics 10.64898/2026.08.19.745830 medRxiv
Top 0.2%
13.2%
Show abstract

Gene model for the ortholog of Insulin-like peptide 3 (Ilp3) in the D. pseudoobscura Apr. 2013 (BCM-HGSC Dpse_3.0/DpseGB3) Genome Assembly (GenBank Accession: GCA_000001765.2) of Drosophila pseudoobscura. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

12
BRAIN CAST: An MRIQC-guided pipeline for age- and sex-specific pediatric brain MRI template construction, validated by downstream structural fidelity

Hu, Y.; Contreras-Vidal, J. L.

2026-08-07 neuroscience 10.64898/2026.08.02.742256 medRxiv
Top 0.2%
13.1%
Show abstract

Pediatric neuroimaging needs age- and sex-appropriate references, yet existing atlases span broad age ranges that blur development or lack sex specificity. We present BRAIN CAST: 28 year-by-year, sex-specific brain MRI templates covering ages 5-18, built from 1,272 quality-screened children in the Healthy Brain Network by an MRIQC-guided pipeline combining reduced-strength denoising, cerebrospinal-fluid-anchored intensity normalization, deep-learning skull stripping and iterative groupwise diffeomorphic registration. We evaluate templates not by image sharpness, which is not comparable across intensity conventions, but by the structural bias they induce downstream. Held-out children align to their matched template with sub-voxel gray-white interface error (1.1 mm); on a direction-symmetric surface-distance metric BRAIN CAST matches the best single-template reference and outperforms an age-specific pediatric atlas in 189 of 189 subjects. Female cortex is fit measurably better by female than by male templates, an effect no sex-neutral reference can provide. Templates, tissue-probability maps and the containerized pipeline are released.

13
Gene model for the ortholog of DENR in Drosophila pseudoobscura

Lawson, M. E.; Sanow, K.; Fratian, M.; Matura, M.; Scanlon, R.; Richard, M.; Nakhla, M.; Rele, C. P.; Thompson, J. S.; Findlay, G. D.; O'Rourke, K. S.

2026-08-11 genomics 10.64898/2026.08.11.744233 medRxiv
Top 0.2%
12.8%
Show abstract

Gene model for the ortholog of Density regulated protein (DENR) in the Apr. 2013 (BCM-HGSC Dpse_3.0/DpseGB3) Genome Assembly (GenBank Accession: GCA_000001765.2) of Drosophila pseudoobscura. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

14
Global vascular plants reveal persistent gaps across taxa and ecoregions

Maciel, E. A.

2026-08-28 evolutionary biology 10.64898/2026.08.24.746674 medRxiv
Top 0.2%
12.7%
Show abstract

Biodiversity aggregators such as GBIF provide unprecedented access to global biodiversity data, yet their representativeness remains uneven across space and taxa. This study examined the spatial and taxonomic structure of global vascular plant data available on GBIF. Six filters were applied to the GBIF vascular plant dataset, resulting in the removal of 54% of all records. Together, the filters explained more than 90% of the identified spatial issues, with duplicate and missing coordinates accounting for most of the variation. A higher number of occurrence records was associated with a greater number of spatial issues. Record distributions became progressively more even at finer taxonomic levels, from orders to species. The time series of occurrences for species, genera, and families increased sharply after 1800 and continued to rise, with no apparent stabilisation. Of the 824 ecoregions covered, 73 accounted for 72% of all occurrence records. These ecoregions spanned all continents but were strongly concentrated in Europe, followed by North America and Oceania. The analyses reveal four key patterns: (1) data volume is positively associated with spatial issues; (2) a small number of taxa account for a large proportion of records, whereas many are represented by relatively few; (3) occurrence data aggregated by GBIF have increased continuously since 1800; and (4) record coverage remains highly uneven across the world's ecoregions. These results highlight the substantial contribution of biodiversity data aggregators to expanding access to biological information while demonstrating the persistent spatial and taxonomic biases that shape their contents. Such biases should be explicitly considered when assessing data completeness and quality and when using aggregated occurrence records to infer global biodiversity patterns.

15
Making Accelerating Medicines Partnership Data Findable and Interoperable through a Common Data Model: Extending OMOP for Multi-Source Multimodal Data

Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.

2026-09-02 genetic and genomic medicine 10.64898/2026.08.31.26361831 medRxiv
Top 0.2%
11.8%
Show abstract

SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.

16
10.5 Tesla High-Resolution Macaque Brain MRI for In vivo and Ex vivo Connectivity Studies

Warrington, S.; Selim, M. K.; Tendler, B. C.; Moeller, S.; Farooq, H.; Wu, W.; Pisharady, P. K.; Adriany, G.; Auerbach, E. J.; Folloni, D.; Bratch, A.; Manea, A. M.; Grafft, T.; Jungst, S.; Harel, N.; Waks, M.; Pestilli, F.; Yacoub, E.; Lenglet, C.; Ugurbil, K.; Heilbronner, S. R.; Miller, K. L.; Jbabdi, S.; Zimmermann, J.; Sotiropoulos, S. N.

2026-08-26 neuroscience 10.64898/2025.12.22.695917 medRxiv
Top 0.2%
11.4%
Show abstract

Mapping brain connectivity in primates remains a major challenge due to difficulties in resolving microscopic white matter architecture, while maintaining whole-brain coverage. Increasing imaging spatial resolution is key for disambiguating fibre configurations within smaller anatomical volumes. Here, we present novel developments that allow high-resolution diffusion MRI of the macaque brain using one of the world's highest-field human MRI scanners operating at 10.5 Tesla, allowing both in vivo and ex vivo macaque brain imaging. Our approach achieves very high spatial resolution across both tissue states, (up to 580 m)3 in vivo and (300 m)3 ex vivo, with diffusion weighting up to b = 6000 s/mm2. We detail methodological advances in data acquisition, image reconstruction, processing and whole-brain tractography that overcome critical challenges associated with ultra-high-field imaging. This work establishes a new framework for high-resolution in vivo and ex vivo neuroimaging of the NHP brain at 10.5 T using a human bore scanner, paving the way for subsequent analyses of brain connectivity across species and tissue states at unprecedented detail. The dataset, along with all processing pipelines, containerised workflows, and reusable web services, is openly shared to support reproducibility and future integration with microscopy for studying white matter microstructure and connections at the mesoscale.

17
Most published human disease RNA-seq cannot be uniformly reanalyzed: a population-scale audit of reprocessability in the Gene Expression Omnibus

Murad, A. B.

2026-08-11 bioinformatics 10.64898/2026.08.05.742891 medRxiv
Top 0.2%
11.0%
Show abstract

BackgroundReuse of archived transcriptomic data underpins a large and growing share of published genomics. Because differences in upstream processing confound cross-study comparison, uniform reprocessing compendia -- recount3, ARCHS4, DEE2, refine.bio, Expression Atlas -- are widely treated as the remedy, and their availability is routinely assumed at the point of study design. Whether that remedy is actually obtainable for the population of published disease RNA-seq has not been measured. Prior audits have characterised metadata completeness and deposition rates, but none has quantified, across the published population, what fraction of studies can be uniformly reprocessed or where in the path from publication to comparable counts that capability is lost. ResultsWe enumerated 1,124 MeSH disease descriptors exhaustively, retrieved 16,820 human RNA-seq series from the Gene Expression Omnibus, and audited the 3,631 bulk, Illumina-platform series of at least 25 samples under three independently pre-registered, tool-enforced analysis plans. Raw reads were publicly available for 94.1% of series under a dual-route evidence standard, but only 46.7% appeared in any uniform reprocessing compendium (bounds 46.7-57.2%) and only 27.3% were usably covered at a 90% run threshold (bounds 27.3-38.8%). Of the 3,418 series whose reads are public, 991 were usably covered, leaving 71.0% of read-public series reprocessed by nothing usable. Presence overstated usability: DEE2 was present for 30.2% of series but usable for 3.4%. Design attrition was independent and severe -- 24.7% met bulk primary-tissue case-control criteria, 7.4% additionally reached a minimum replication threshold counted on sample accessions, and 4.1% did so counted on distinct donors. Among the 199 series where donor identity resolves, 4.1% pass the replication criterion on donors against 15.1% on accessions, a 3.73-fold difference; across the census frame, accessions exceeded distinct donors by 2.65-fold (Manski bounds 1.09-10.77, Imbens-Manski 95% CI 1.07-11.88). Independently, 36.4% of series-to-disease attributions produced by a conventional keyword query were refuted by the curated MeSH headings of the series own linked publication. ConclusionsUniform reanalysis of published human disease RNA-seq is unavailable for most studies in the population audited, and the binding constraint is usable coverage rather than deposition of raw reads. The loss occurs at several independent layers with different remedies, and the coverage layer -- unlike the others -- is one that resource maintainers can act on. Automated retrieval further over-counts eligible studies, both by admitting designs outside scope and by assigning studies to diseases their publications do not support.

18
EegFun.jl: A Julia Package Tutorial for EEG Analysis

Dudschig, C.; Sonntag, S.; Mackenzie, I. G.

2026-08-12 neuroscience 10.64898/2026.08.11.744163 medRxiv
Top 0.3%
8.0%
Show abstract

EegFun.jl is an open-source package for electroencephalography (EEG) analysis implemented in the Julia programming language. EegFun.jl provides a flexible framework for EEG research, covering data import from standard file formats, filtering and re-referencing, Independent Component Analysis (ICA) for artifact detection/correction, epoch extraction, and ERP averaging and visualisation. The Julia language provides the readability of a high-level scripting environment together with execution speeds comparable to compiled code. EegFun.jl combines interactive data visualization with high-performance execution, making large-scale analyses both efficient and easy. Here, we provide a brief overview and introductory tutorial of the core stages of the EEG analysis workflow to illustrate the packages capabilities. The package is freely available under the MIT license.

19
MAESTRO: A Public, Generalizable Model for Stroke Lesion Segmentation from T1 MRI Across the Recovery Continuum

Khan, M. H.; Marin-Pardo, O.; Chakraborty, S.; Lee, K.; Lee, S. Y.; Raman, N.; Iglesias, J. E.; Liew, S.-L.

2026-08-25 neurology 10.64898/2026.08.22.26361044 medRxiv
Top 0.3%
8.0%
Show abstract

Accurate stroke lesion segmentation is essential for large-scale neuroimaging studies, yet manual delineation remains labor-intensive, and existing automated methods often struggle to generalize across imaging protocols and stages of recovery. We developed MAESTRO, a deep learning framework for automated lesion segmentation across the stroke recovery continuum using T1-weighted (T1) MRI alone. We hypothesized that combining a transformer-based architecture with an image augmentation strategy would improve segmentation accuracy and robustness under heterogeneous imaging conditions. T1 MRI scans and expert-traced lesion masks from 955 stroke participants across 33 international cohorts were used to train and evaluate MAESTRO within the open-source nnU-Net framework. Performance was evaluated on a held-out test set using spatial and volumetric agreement metrics. An exploratory human-in-the-loop (HITL) evaluation compared correction of MAESTRO-generated segmentations with manual tracing from scratch. MAESTRO achieved the strongest performance across several evaluated model configurations, providing the most accurate lesion localization and lesion volume estimates (median Dice = 0.686; Pearson r = 0.861; ICC = 0.792). Segmentation performance was sensitive to lesion size and stroke chronicity but remained robust across diverse imaging conditions. Additionally, using a HITL workflow to correct MAESTRO segmentations reduced annotation time by 47.4% compared to manual tracing while improving accuracy relative to both automated and manual workflows. MAESTRO is publicly available to enable robust, automated stroke lesion segmentation from T1 MRI. When combined with human review and correction, MAESTRO offers a practical approach for generating standardized, high-quality lesion annotations, helping reduce a major practical barrier to large-scale stroke imaging studies.

20
A framework for quality assurance in human intracranial electrophysiology

Herz, N.; Cao, R.; Qiu, S.

2026-08-12 neuroscience 10.64898/2026.08.06.743130 medRxiv
Top 0.3%
8.0%
Show abstract

Intracranial electroencephalography (iEEG) provides an unprecedented opportunity to directly record neural activity and causally perturb the human brain through electrical stimulation. Yet, the increasingly collaborative nature and complexity of modern iEEG studies pose substantial challenges for experimental control, data quality, and standardization. Unlike most experimental modalities, human iEEG data are acquired within dynamic clinical environments, where patient condition, recording quality, hardware configuration, and experimental protocols may vary across recording sessions and collaborating sites. The resulting heterogeneity creates opportunities for technical and procedural failures that often remain undetected until downstream analyses, when corrective action is no longer possible. Here, we present a framework for standardized session-level quality assurance in human iEEG research and provide an open-source implementation compatible with Brain Imaging Data Structure (BIDS)-organized datasets. The framework defines four complementary domains of quality assessment crucial for human iEEG studies: protocol fidelity, behavioral integrity, stimulation validation, and signal quality. These domains integrate electrophysiological recordings, behavioral event logs, and stimulation metadata to verify data completeness, confirm participant engagement, validate stimulation delivery, and identify potentially compromised recording channels. Automated quality metrics and standardized diagnostic visualizations are generated following each testing session, enabling rapid identification of technical and procedural failures while corrective action is still possible. By providing a standardized approach to session-level quality assurance, the framework improves data integrity, enhances reproducibility, facilitates analyst training, and supports harmonized data collection across laboratories and clinical sites.